Skip to main content

4.2 Why Activation Function?

Without an activation function, stacking layers isnt multiple functions, it can be simplified into a single linear function, we dont want that so we are sticking relu in the middle. ReLU introduces non-linearity.

A neuron can therefore calculate:

z=Wx+bz=Wx+b
ReLU stands for Rectified Linear Unit.SigmoidTanh
ReLU(x)=max⁡(0,x)\boxed{ReLU(x)=\max(0,x)}σ(x)=11+e−x\boxed{\sigma(x)=\frac{1}{1+e^{-x}}}tanh(x)=ex−e−xex+e−x\boxed{tanh(x)=\frac{e^x-e^{-x}}{e^x+e^{-x}}}
ReLU(5)=5ReLU(5)=5
ReLU(−3)=0ReLU(-3)=0
σ(0)=0.5\sigma(0)=0.5
σ(10)≈1\sigma(10)\approx1
tanh(−10)≈−1tanh(-10)\approx-1
tanh(10)≈1tanh(10)\approx1
return max(0, x)return 1 / (1 + math.exp(-x))return math.tanh(x)